Chinese Chunking and Consistency Checking Using Rule-Based Method

نویسندگان

  • Jiao-Li Lu
  • Jia-Heng Zheng
  • Hong-Ye Tan
  • Jian Sun
چکیده

This paper presents a rule-based chunking approach. Rule-based method does well in analyzing the structure of natural language. In order to avoid the confliction of the rules, we extract a small scale chunking rule set for chunking first. Then we define more rules to check and correct the inconsistency phenomena. We also adopt man-machine interaction method to solve some special language phenomena. Experimental results show that our approach achieves high accuracy.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

A Grammar Checking System for Punjabi

This article provides description about the grammar checking system developed for detecting various grammatical errors in Punjabi texts. This system utilizes a fullform lexicon for morphological analysis, and applies rule-based approaches for part-of-speech tagging and phrase chunking. The system follows a novel approach of performing agreement checks at phrase and clause levels using the gramm...

متن کامل

A Study on Consistency Checking Method of Part-Of-Speech Tagging for Chinese Corpora

Ensuring consistency of Part-Of-Speech (POS) tagging plays an important role in the construction of high-quality Chinese corpora. After having analyzed the POS tagging of multi-category words in large-scale corpora, we propose a novel classification-based consistency checking method of POS tagging in this paper. Our method builds a vector model of the context of multi-category words along with ...

متن کامل

Chunking Using Conditional Random Fields in Korean Texts

We present a method of chunking in Korean texts using conditional random fields (CRFs), a recently introduced probabilistic model for labeling and segmenting sequence of data. In agglutinative languages such as Korean and Japanese, a rule-based chunking method is predominantly used for its simplicity and efficiency. A hybrid of a rule-based and machine learning method was also proposed to handl...

متن کامل

Chinese Chunking Based on Maximum Entropy Markov Models

This paper presents a new Chinese chunking method based on maximum entropy Markov models. We firstly present two types of Chinese chunking specifications and data sets, based on which the chunking models are applied. Then we describe the hidden Markov chunking model and maximum entropy chunking model. Based on our analysis of the two models, we propose a maximum entropy Markov chunking model th...

متن کامل

Automatic Rule Acquisition for Chinese Intra-chunk Relations

Multiword chunking is defined as a task to automatically analyze the external function and internal structure of the multiword chunk(MWC) in a sentence. To deal with this problem, we proposed a rule acquisition algorithm to automatically learn a chunk rule base, under the support of a large scale annotated corpus and a lexical knowledge base. We also proposed an expectation precision index to o...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:
  • JCIT

دوره 5  شماره 

صفحات  -

تاریخ انتشار 2010